What AI visibility actually measures, and what it does not
AI visibility is the share of AI answers you appear in across a fixed set of prompts, measured separately for each assistant. That is the entire definition. It is not your Google ranking. It is not your organic traffic. It is not one blended score, and it is not a proxy for revenue. If you cannot name the twenty prompts you measured, you do not have a measurement.
The unit of measurement is the prompt, not the keyword
A keyword is a fragment. A prompt is a full question with context attached, and the context changes the answer. "AI marketing agency" is a keyword. "Which agencies specialize in AI search visibility for B2B SaaS companies selling into the US?" is a prompt, and it returns a short named list instead of a page of blue links.
Assistants do not rank. They compose. Every answer is assembled at request time, and your presence is close to binary: named, or absent. That is why the measurement fixes the question and repeats it. A rank tracker watches a position move on a stable list. A visibility check asks the same question repeatedly and counts how often you appear.
How to build a 20-prompt set that holds up
Twenty is the practical number: big enough that one lucky appearance does not swing the percentage, small enough that a person will actually run it monthly.
-
Build it across four intents, five prompts each. Category questions (what is answer engine optimization).
-
Selection questions (best agencies for X in the Y market).
-
Comparison questions (X versus Z for a mid-market team).
-
Problem questions (our traffic dropped after a core update, what should we check).
Write them once, then freeze the wording. Run the identical text against each assistant and log four things: whether your brand appeared, whether one of your URLs was cited, which competitors appeared beside you, and the date. Visibility is appearances divided by prompts, reported per assistant. Fifteen percent on one assistant and forty percent on another is a normal result. A single averaged number hides the only thing worth knowing, which is where you are invisible.
Why a fixed set beats checking when you feel like it
Ad hoc checking produces anecdotes. You ask a question, see yourself, feel good, and never ask again. Or you use slightly different wording next month, get a different answer, and cannot tell whether the model changed or you did. Model output varies between runs on identical input, which is exactly why the set must be frozen and repeated: you are measuring a rate, not an event. Without a fixed set, every reading is a coin flip you are reading as a trend.
What we found when we ran this on our own site
We audited growthym.com in September 2026 and crawled 145 URLs. The picture was ugly in ordinary ways: 42 service pages across three overlapping page systems, 31 title tags past 60 characters, a mobile Time to Interactive of 12 seconds caused by a developer CSS tool shipped to production, and domain authority of 14 against 59 organic visits a month.
Then the referral data showed something worse. ChatGPT was sending clicks to a Growthym URL that returned a 404, logged as utm_source=chatgpt.com. An assistant was citing a page that did not exist. We were visible and broken at once, and no ranking report surfaces that. That check is now the first thing Growthym runs on any site, before a prompt gets scored.
What the number does not tell you
Visibility is not traffic. Most AI answers resolve without a click, and a high appearance rate can sit beside flat sessions. It is not demand either: twenty prompts you wrote yourself sample the market's questions rather than count them. It does not capture how you were described: being listed fourth in a neutral sentence is not the same as being named as the recommended option.
Treat it as a leading indicator with a narrow job: it tells you whether assistants know you exist, for questions you chose. That is enough to act on, which is why our AI search visibility work starts with the prompt set rather than a content calendar.
Write your twenty prompts this week. Run them once against two assistants, save the raw answer text rather than a summary, and file it under today's date. That file is your baseline, and you cannot create it retroactively.
What's Your Reaction?





